Back

Journal of Genetics and Genomics

Elsevier BV

Preprints posted in the last 30 days, ranked by how well they match Journal of Genetics and Genomics's content profile, based on 38 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
A Korean pangenome reference of 14 healthy individuals supports structural variant analysis in disease genomes

Shin, D.-H.; Jeon, J.; Joe, S.; Jeon, Y.; Yang, J. O.; Bhak, J.; Baek, S. A.; Byun, G.; Shin, E.-S.; Kwon, Y.; Choi, H.-J.; Kim, J.-H.; Haam, K.; Yoo, J.; Song, K. J.; Mok, J.; Jeon, S.; Jeong, H.; Bhak, J.

2026-07-09 genetic and genomic medicine 10.64898/2026.07.06.26357367 medRxiv
Top 0.1%
2.7%
Show abstract

Here, we present the first graph-based Korean Pangenome Reference (K-PanRef), constructed from 14 healthy Korean individuals. K-PanRef comprises 13 high-quality diploid Korean genome assemblies (mean QV ~62.0) and KOREF1-G-TTAGGA, the first complete Korean reference genome. Integration of these assemblies generated a ~3.2-Gb pangenome graph containing ~39.3 million nodes and ~53.8 million edges, with the accumulation of common sequences (frequency [≥]10%) reaching a plateau. Additionally, K-PanRef contains ~4.3 million Korean-specific small variants and ~76.0 thousand Korean-specific SVs absent from the Chinese and human pangenome references, improving the representation of Korean genetic diversity relative to these references. To evaluate its utility for short-read-based SV analysis, we genotyped 75 whole-genome sequencing (WGS) samples, including 15 patients with early-onset myocardial infarction (MI). Although constructed entirely from healthy genomes, K-PanRef supported the identification of putative disease-relevant SVs in this exploratory application. K-PanRef-based genotyping identified ~95.6 thousand small variants and 820 SVs observed only in the early-onset MI samples. Among the early-onset MI-group SVs, 491 were absent from public databases, suggesting that they may represent previously unrecognized candidate variants related to early-onset MI. Of these, 164 SVs overlapped 134 genes, of which 89 had reported associations with 42 cardiovascular diseases or traits, including eight genes previously linked to MI. Together, these results establish K-PanRef as a valuable resource for representing Korean genetic diversity and enabling more comprehensive discovery of population-specific and novel putative disease-relevant variants from short-read sequencing data.

2
In Vivo Spatial Transcriptomics for Bleeding-free Profiling Human Internal Organs

Sun, H.; Guo, F.; Zhao, X.; Wan, Y.; Zhang, X.; Sun, J.; He, X.; Gai, B.; Xiong, C.; Ma, Y.; Qu, J.; Li, P.; Gao, F.; Zhao, X.; Ji, X.; Yang, Z.; Mak, L.-Y.; Yap, Y. H.; Ke, J.; Shi, P.

2026-07-09 genetic and genomic medicine 10.64898/2026.07.06.26357355 medRxiv
Top 0.5%
1.1%
Show abstract

Despite the significant technical advancement in spatial transcriptomics, its clinical usage is largely untapped. Here, we develop an integrated system, ENDO-Genome, for minimally invasive in-body transcript sampling to facilitate live spatial transcriptomic analysis of human internal organs. This is achieved by integrating a nanoarrayed biochip with existing endoscope to perform pressure-sensor-calibrated "Touch & Go" RNA extraction directly from human internal organs, including the highly vascularized liver or kidney, without the need for tissue biopsy, voiding any bleeding risks. By a demonstration using gastrointestinal endoscopy, multiplexed landscape of 55 mRNA transcripts was obtained from multiple locations of human intestinal tract via a 5-minute operation in routine examinations. Benefiting from a sequencing-free approach, each assay costs less than 10 US dollars. For the clinical study involving 15 Crohn' s disease (CD) patients, no complication case was reported out of 47 ENDO-Genome operations, showcasing the gentle deposition and excellent safety of the technique. The live spatial transcriptomics provides direct in vivo pictures of the heterogenous spatial transcriptional programs underlying CD pathological response at different intestinal locations, revealing distinct ileal phenotypes. This is manifested by unique microscale scattering of inflammation gene clusters, along with the discovery of a tissue-specific cooperative mechanisms between inflammation and RNA methylation regulations at single- or multi-cell scales.

3
Whole-genome duplication underlies conserved sexually biased expression of meiotic cohesin genes unique to the teleost fish lineage

Niwa, T.;Kikuchi, M.;Tanaka, M.

2026-06-27 Developmental Biology 10.64898/2026.06.26.731870 medRxiv
Top 0.6%
1.0%
Show abstract

Meiosis is a fundamental process in producing both sperm and eggs, yet recombination landscapes often exhibit sexual differences, known as heterochiasmy. Since meiotic proteins are generally expressed in both sexes, the molecular mechanism driving heterochiasmy remains elusive. The -kleisin subunit gene of meiotic cohesin, Rec8, is expressed bisexually in mammals, while its putative teleost ortholog, rec8a, is expressed in a female-biased manner, presumably due to the presence of its paralog originating from the teleost-specific whole-genome duplication (TGD). Here, we elucidated the evolutionary history and expression dynamics of -kleisin genes across teleost lineages. Through comprehensive phylogenetic and synteny analyses, we revealed that major teleost lineages retain two copies of rec8 and rad21, with rec8 loci experiencing drastic chromosomal rearrangements immediately after the TGD. Using in situ hybridization and single-cell transcriptome data in medaka and zebrafish, we demonstrated a conserved sexually biased expression pattern: rec8a is predominantly female-biased, whereas rec8b exhibits male-biased expression during gametogenesis. Furthermore, comparative epigenetic analyses revealed that the conserved sexually biased expression is driven by lineage-specific cis-regulatory elements, rather than conserved ones. Motif analyses imply that regulatory rewiring by transcription factors, including foxl2l in particular, might have played a crucial role in the establishment and maintenance of this paralog divergence. Our findings highlight how whole-genome duplication and subsequent genomic and epigenetic rewiring subdivided the bisexual function of rec8, offering insights into sexually distinct meiotic regulation. HighlightsO_LITeleosts possess a unique -kleisin repertoire originating from the TGD. C_LIO_LITeleost rec8 paralogs exhibit conserved sex-biased expression during meiosis. C_LIO_LIDrastic genomic rearrangements after the duplication rewired the teleost rec8 loci. C_LIO_LIThe conserved expression pattern is governed by lineage-specific CREs. C_LIO_LIThose CREs harbor similar types of TFBSs such as Fox-family TFs. C_LI Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=94 SRC="FIGDIR/small/731870v1_ufig1.gif" ALT="Figure 1"> View larger version (22K): org.highwire.dtl.DTLVardef@c8c84dorg.highwire.dtl.DTLVardef@1d65668org.highwire.dtl.DTLVardef@c2d732org.highwire.dtl.DTLVardef@1be54a2_HPS_FORMAT_FIGEXP M_FIG C_FIG

4
Chromosome-level genome assembly of Calotes wangi with dynamic colour variation

Qiu, X.; Wang, Y.; Wen, J.; Chen, Y.; Zhao, L.; Jian, J.; Yang, W.

2026-07-10 evolutionary biology 10.64898/2026.07.07.736949 medRxiv
Top 0.8%
0.8%
Show abstract

The Wangs garden lizard, Calotes wangi, is a widely distributed agamid species in Southern China and Northern Vietnam and exhibits pronounced colour variation and rapid body colour change. Despite increasing interest in the genomic basis of colour variation, chromosome-level genomic resources remain limited in agamid lizards. Here, we generated a chromosome-level reference genome of C. wangi using PacBio HiFi sequencing and Hi-C scaffolding. The final genome assembly was approximately 1.66 Gb in size and comprised 6 macrochromosomes and 11 microchromosomes, with a contig N50 of 110.09 Mb and 98.9% complete BUSCO genes. A total of 20,442 protein-coding genes were annotated. Comparative genomic analyses identified 297 significantly expanded gene families, with enriched functions associated with steroid metabolism, chromatin regulation, and epigenetic processes. This high-quality genome assembly provides an important genomic resource for future studies of colour variation, phenotypic plasticity, and evolutionary diversification in agamid lizards.

5
SPEAK: Spatial Prompting with Expert Aligned Knowledge for Tissue Domain Identification in Spatial Transcriptomics

Wei, H.; Luo, X.; Yu, H.; Liang, J.; Yang, L.; Sauler, M.; Kaminski, N.; Popa, A.; Yan, X.

2026-06-26 bioinformatics 10.64898/2026.06.22.733750 medRxiv
Top 0.9%
0.6%
Show abstract

Spatially resolved transcriptomic (SRT) data requires spatial domain identification to enable tissue microenvironment-specific downstream analyses. Here we present SPEAK (Spatial Prompting with Expert-Aligned Knowledge), a large language model (LLM)-based method to identify spatial domains from SRT data by taking advantage of the prior knowledge from both LLM and human experts. SPEAK constructs a spatial context prompt for each cell/spot based on cell types and marker genes of its neighboring cells, enabling zero-shot inference, expert-guided fine-tuning, and prototype updating through two-stage prompting. Applications to STARmap, Visium, MERFISH and Xenium datasets showed advantages of SPEAK over existing spatial domain identification methods in domain prediction accuracy, robustness to limited prior knowledge, biological interpretability, and capacity for efficient expert-guided fine-tuning with generalizability to other tissue sections.

6
Mapping non-coding functional elements in allotetraploid Cyprinus carpio embryo development reveals subgenome variation of transcription regulation

Jimenez-Gonzalez, A.; Madrero Pardo, A.; Hadzhiev, Y.; Blasweiler, A.; Zunar, B.; Csenki-Bakos, Z.; Muller, T.; Megens, H.-J.; Wiegertjes, G. F.; Lenhard, B.; Baranasic, D.; Mueller, F.

2026-06-26 genomics 10.64898/2026.06.22.733679 medRxiv
Top 1.0%
0.6%
Show abstract

Common carp (Cyprinus carpio) is an important freshwater species for ornamental and aquaculture purposes, and a key cyprinid model for studying allotetraploidy. Its two chromosomally-separated subgenomes show distinct gene expression profiles, but how their regulatory landscapes control gene expression dynamics during development remains unknown. We generated a regulatory atlas by combining transcriptomes across 12 developmental stages with chromatin accessibility maps, transcription start sites and gene regulation-associated histone post-translational modifications. Subgenome-specific annotation and comparison of 254,276 developmental regulatory elements (PADREs) revealed that regulatory subgenome divergence is most prominent during early development, converging toward the phylotypic period, mirroring expression convergence between subgenomes at the same stages. This dynamic was driven by enhancers, while promoters maintained a more stable subgenome bias, extending the hourglass model of developmental constraint to allotetraploid subgenome regulation. Subgenome-specific enhancers were preferentially retained in subgenome B, whereas subgenome A shifted toward homeologous enhancer activity near the phylotypic stage, indicating directional regulatory divergence between subgenomes. Comparison with zebrafish revealed high concordance with sequence conservation and that subgenome B retained more ancestral cyprinid regulatory elements than subgenome A. This developmental regulatory atlas provides a foundational resource for investigating cis-regulatory evolution following the fourth round of vertebrate genome duplication.

7
A cis-regulatory variant in ASIP causes gray coat color in the donkey

li, y.; Liu, Y.; wu, j.; liu, s.; lin, x.; guo, k.; yang, t.; feng, m.; zhang, h.; wang, x.; xing, w.; qian, s.; yang, r.; zhao, c.

2026-06-28 genomics 10.64898/2026.06.23.733953 medRxiv
Top 1%
0.5%
Show abstract

BackgroundGray is one of the relatively rare coat colors in donkeys. The Hetian Gray donkey is a distinctive indigenous breed from the Xinjiang Uygur Autonomous Region of Northwestern China, characterized by progressive hair depigmentation with aging while retaining dark skin pigmentation. However, the genetic basis underlying this unique gray coat color phenotype remains unclear. ResultsTo elucidate the genetic basis, we conducted whole-genome resequencing of Gray and non-Gray donkeys. Genome-wide selection signature analyses identified a candidate region on chromosome 15. Subsequent fine-mapping using mass spectrometry-based genotyping of 42 loci refined the candidate interval and revealed a SNP within intron 2 of the ASIP gene, located in a genomic fragment with highly similar sequences, showing complete association with the gray coat color. Association analysis in an expanded population further confirmed a strong correlation between this variant and the gray phenotype. Gene expression analyses also supported the role of ASIP in regulating pigmentation in donkeys. ConclusionsThese findings identify a genetic determinant of gray coat color in donkeys and provide new insights into the molecular mechanisms underlying age-related depigmentation in domestic animals.

8
Conserved Genic Composition Unravels Rapid Karyotype Evolution and Polyploidization across Plants

Gu, J.; Chen, W.; Li, D.; Tang, H.; Li, X.

2026-07-02 evolutionary biology 10.64898/2026.06.29.735181 medRxiv
Top 1%
0.5%
Show abstract

Chromosomal variation underlies species evolution, but reconstructing its large-scale dynamics remains challenging, obscuring its adaptive significance. Here, we introduce GouMang, a framework that mines conserved genic section compositions across diverse species to trace karyotype evolution. In grasses, applied to 818 highly varied chromosomes spanning eight subfamilies, GouMang resolved a shared karyotype evolution path of 9-to-18 ({rho} whole genome duplication, {rho}WGD)-to-12 chromosomes, followed by lineage specific rearrangements or WGDs. Genes retained from the early {rho}WGD are linked to cold/light adaptation, supporting a key biomass expansion event that impacted subsequent global ecological pattern and human agricultural civilization, in which K-Pg global cooling and subsequent forest degradation drove early understory grasses to sun plants. Parallel analysis in Brassicaceae reconstructed karyotype evolution as well as {beta}WGD which unlinked to cold/light adaptation, reflecting a divergent biogeographic history compared to grasses. Together, GouMang depicts a widespread plant evolutionary pattern where karyotype constantly diversified with WGDs recurrently fueling adaptation.

9
Self-Supervised Behavioral Representations Across the Life Course: A Killifish Case Study

Chang, J.-C.; Komatsu, T. S.; Onami, S.

2026-06-29 animal behavior and cognition 10.64898/2026.06.23.733896 medRxiv
Top 1%
0.5%
Show abstract

Self-supervised foundation models of aging are increasingly built from longitudinal data (biobanks, electronic health records, wearables) that is inherently incomplete: no individual is followed across a whole lifetime, and how much of each life is captured varies widely. This raises two linked questions: is it worth modeling an individual's whole life course rather than its current state, and can such a model be built from brief, fragmentary records? No human cohort can settle them, because none offers a complete life to compare against. We turn to the African turquoise killifish (Nothobranchius furzeri), tracked from youth to natural death in publicly released recordings, as a controlled testbed: its complete lifespans provide the full-life reference that human data lacks. On these data we build LifeMAE, a two-stage selfsupervised model: a day encoder that summarizes each day of behavior, then a life-course encoder over the trajectory of those daily summaries. We find that the day encoder alone is already strong: from a single day of behavior it predicts chronological age, separates long- from short-lived individuals (coarsely), and flags nearness to death. Adding the life-course encoder improves on none of the three; each is matched by trivially aggregating the day-level predictions (a smoother for age, an early-life average for lifespan). Near-term mortality seems the exception, where the whole-life model looks far better (AUROC 0.81 to 0.91), but the gain is not behavioral: it reflects where each day falls within the observation window (a cue supplied by the model's encoding of time), and a single-day model given that cue closes the gap at any observation length. For these traits, an individual's place in its life course is legible from a single day: the trajectory stage is unnecessary, and the record it needs is as short as one day, the finest grain our day-level setup resolves. For characterizing a cohort, this favors observing many individuals briefly over tracking a few for long. The result joins a growing body of work in which deep and foundation models, fairly benchmarked, fail to beat deliberately simple baselines. We add a concrete mechanism for the over-optimism: a model's encoding of time can leak the very quantity it predicts, which backwardlooking evaluation mistakes for learned biology, so only evaluation fixed to the moment of prediction is trustworthy.

10
Structural Variant Imputation in Samoans Using a Population-Specific Reference Panel

Spor, L. M.; Liau, E. M.; Sanchis-Juan, A.; Silva, A. N.; Cheng, H.; Naseri, T.; Reupena, M. S.; Viali, S.; Tuitele, J.; Kershaw, E. E.; Deka, R. D.; Hawley, N. L.; McGarvey, S. T.; Weeks, D. E.; Carlson, J. C.; Brand, H.; Minster, R. L.

2026-07-10 genetic and genomic medicine 10.64898/2026.07.07.26357461 medRxiv
Top 1%
0.5%
Show abstract

Structural variants (SVs) are often excluded from genetic research because they are difficult to call, but they can have substantial effects on phenotypic traits. SVs have not previously been characterized in Samoans, an understudied population with a high burden of complex diseases. Using short-read whole genome sequencing data, we called SVs in 1,276 Samoans and created a Samoan-specific imputation panel inclusive of both SVs and single nucleotide variants (SNVs), called the Soifua Manuia-SV panel. Using this panel, we imputed SVs and SNVs in 3,611 Samoans with array data, enabling analysis of SV-phenotype associations in a sample of 4,887 Samoan participants. We evaluated imputation performance in Samoans against two other reference panels: (i) an SNV-only Samoan-specific reference panel, to assess whether SV inclusion impacts SNV imputation, and (ii) an SV and SNV, multi-ancestry reference panel composed of 1000 Genomes participants, which did not include Polynesians, to assess the importance of including the target population in the reference panel. The Soifua Manuia-SV panel substantially outperformed the multi-ancestry SV and SNV panel, yielding 5.5 million more high-quality (r2[≥]0.8) variants, including over 8,000 more high-quality SVs. SNV imputation based on the two Samoan-specific panels performed similarly overall, suggesting that SV inclusion does not strongly impact SNV imputation quality. This work highlights the importance of population representation for accurate imputation.

11
Suspension-Based Human Esophagoids Recapitulate WNT2B-Dependent Regulation of Esophageal Basal Progenitors

Etzioni, N.; Frum, T.; Johnson, K.; Alvarez-Maldonado, A. P.; Yllescas-Lopez, H. M.; Bayer, D. E.; Xiao, Z.; Eiken, M. K.; Loebel, C.; Wu, J. H.; Tsai, Y.-H.; Wu, A.; Zhang, C. J.; Dame, M. K.; Gunuguntla, B.; Cuttitta, A. J.; Ho, H.; Tigani, D. J.; Sexton, J.; Dasuri, V. S.; Makogonov, N.; OConnell, A. E.; Spence, J. R.; Torres, D. F.

2026-07-08 developmental biology 10.64898/2026.06.23.733451 medRxiv
Top 1%
0.5%
Show abstract

Background & AimsThe human esophagus undergoes a tightly regulated developmental program, transitioning from a simple columnar epithelium in early development to a mature stratified squamous tissue essential for adult barrier function. Here, we constructed a developmental cell atlas spanning early development to adulthood and leveraged it to generate physiologically relevant in vitro models. MethodsWe utilized single-cell RNA sequencing and spatial multiplex proteomics of human esophageal tissue from early development through adulthood. We established a feeder-supported 2D culture system and a Matrigel-free, suspension-based 3D esophagoid model in a 96-well format. To interrogate WNT2B function, we analyzed patient tissue harboring WNT2B loss-of-function mutations and performed WNT inhibition in esophagoids. ResultsSequencing profiling identified stage-specific epithelial populations: multiciliated and GPC3 basal cells were unique to early development; KRT14 basal and CRNN luminal cells were adult-specific; and COL17A1, LY6D, and KRT4 populations were shared across stages. Spatially organized WNT2B, KIT, and VWC2 mesenchymal subtypes were identified. The 2D system preserved both epithelial and mesenchymal compartments with transcriptional fidelity. Esophagoids exhibited basal-to-luminal stratification, mesenchymal compartmentalization, and required stromal interactions for formation. WNT2B repressed self-renewal of TP63 basal progenitors and inhibited proliferation, confirmed by pharmacologic inhibition of WNT in the in vitro esophagoids. ConclusionsWe present a stage-resolved atlas of human esophageal development and a scalable esophagoid platform recapitulating esophageal architecture. WNT2B regulates progenitor dynamics by restraining basal cell self-renewal. Esophagoids provide a physiologically relevant system for modeling esophageal development and disease. Visual Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=139 SRC="FIGDIR/small/733451v1_ufig1.gif" ALT="Figure 1"> View larger version (45K): org.highwire.dtl.DTLVardef@1af5743org.highwire.dtl.DTLVardef@8a061dorg.highwire.dtl.DTLVardef@1975c3forg.highwire.dtl.DTLVardef@292ce9_HPS_FORMAT_FIGEXP M_FIG C_FIG Key Findings and ImplicationsO_LIDevelopmental Atlas: The study presents a comprehensive transcriptional and structural atlas of the human esophageal epithelium, identifying conserved and stage-specific epithelial populations from early development to adulthood. Notably, stage-specific gene expression of multiciliated and GPC3 basal cells were unique to early development, while KRT14 basal and CRNN luminal cells were adult specific, with COL17A1+ (basal), LY6D+ (epibasal), and KRT4+ (middle), shared at all stages. C_LIO_LIMesenchymal Diversity: Spatial and transcriptional profiling revealed distinct mesenchymal subtypes, including WNT2B, KIT, and VWC2 populations, which are spatially organized and contribute to epithelial-mesenchymal signaling. These findings reinforce the role of stromal-epithelial interactions in esophageal development. C_LIO_LI2D Esophagus Cell Culture System: A feeder-supported 2D cell culture system was developed that retains both epithelial and mesenchymal populations, preserving transcriptional fidelity and enabling long-term expansion for mechanistic studies. C_LIO_LI3D Esophagoid Model: A suspension-based 3D organoid system was optimized using a 96-well format, enabling high-throughput generation of esophagoids with robust epithelial stratification and mesenchymal compartmentalization. These organoids recapitulate key features of the human esophagus, including basal-to-luminal organization, and require stromal interactions for formation. C_LIO_LIFunctional Role of WNT2B in esophagus development: Both in vivo and in vitro analyses demonstrated that WNT2B regulates epithelial progenitor dynamics and tissue architecture by repressing self-renewal of basally localized TP63+ cells and inhibiting proliferation. Loss-of-function models and WNT pathway modulation confirmed its role in epithelial-mesenchymal crosstalk and organoid integrity. C_LI

12
Residual Multi-Modal Learning for Pan-Breast-Cancer Drug Response Prediction

Huang, B.; Tasaka, L.; Li, J.; Islam, T.; Zhang, S.

2026-07-08 bioinformatics 10.64898/2026.07.03.736239 medRxiv
Top 1%
0.5%
Show abstract

Predicting drug sensitivity across diverse cancer cell lines remains a fundamental challenge in precision oncology, particularly for data-scarce cell lines where per-cell-line models overfit and lookup-table approaches cannot generalise to unseen biological contexts. We present DL4DR, a Two Tower Residual Late Fusion deep learning model that addresses this challenge through content-based, identity-free genomic conditioning. The Cell Line Tower encodes each cell line as a 3 x 139 x 139 genomic image - encoding gene expression, mutation severity, and copy-number variation as RGB channels - using a convolutional encoder that maps directly from biological content, never from a cell line ID. The Compound Tower combines three complementary molecular representations: D-MPNN graph message passing, ORNN octave convolutional image features, and an ECFP hard-memorization head that preserves activity-cliff resolution. Predictions are composed as a residual sum: f = fhard + {lambda}(zc). fresidual, where the learned gate $\lambda$ modulates how much interaction signal supplements the memorization baseline. Evaluated across 51 breast cancer cell lines(136,342 records), Residual Fusion outperforms the ECFP-Only baseline in 48/51 cell lines (94.1%), with {Delta}R2 > 0.02 in 26/51 (51.0%). On the leave-cell-line-out split - the decisive test of genomic generalisation - the mean {Delta} R2 = 0.016 across all 51 lines demonstrates that the genomic encoder learns transferable biological signal beyond cell line identity. External validation on 601 cell lines across 27 cancer tissue types (CellTiter-Glo dataset; 0 cell line overlap with training) achieves median R2 = 0.627, within the range of the internal random-split performance (R2 = 0.61--0.69), confirming pan-cancer generalisation. GradCAM interpretability on the Cell Line Tower recovers TP53 among the top-five cross-cell-line genomic activators (5/51 cell lines) alongside several uncharacterised candidate genes (e.g.FSIP2, 6/51) - without any prior pathway annotation - providing partial biological validation of the learned representation, while also indicating that a substantial share of the encoder's top-ranked signal corresponds to genes with no current annotation as breast cancer drivers. Code and data are available at https://github.com/bayjuan5/DL4DR.

13
Safeguarding open-weight genomic foundation models through weight locking

Karatzikos, A.; Vasilopoulou, A.; Chan, C.; Mouratidis, I.; Georgakopoulos-Soares, I.

2026-07-10 bioinformatics 10.64898/2026.07.07.736795 medRxiv
Top 1%
0.5%
Show abstract

BackgroundGenomic foundation models can dramatically accelerate biological research by learning general-purpose representations of genomic data that transfer across tasks, enabling researchers to predict variant effects, regulatory elements, and molecular function, among others. To safeguard against potential biosecurity threats and malicious misuse of open-weight models, a common strategy involves excluding human-infecting viral genomes from the models training corpora. This strategy, however, can be easily circumvented by fine-tuning models on abundantly available viral data. Weight-locking with spectral deformation has been proposed as a potential method to prevent fine-tuning of neural networks, but has not been systematically evaluated in biological AI models. MethodsWe applied spectral deformation locking to the Evo-1-8k-base genomic foundation model and evaluated a panel of attack configurations spanning naive fine-tuning, low-rank adaptation (LoRA), a simple inserted-layer bypass baseline, and a white-box singular value decomposition (SVD)-chain factorisation at chain lengths k [isin] {2, 3, 5}. Recovered virological capability was quantified on three Human Virome Understanding Evaluation (HVUE) tasks. ResultsThe lock defended against the naive attacker by either standard pipeline. Naive full fine-tuning under the strong lock drove downstream virological capability significantly below the pretrained baseline on pathogenicity and host tropism, converting the attack into a capability loss rather than a gain, while naive low-rank adaptation neither moved held-out perplexity (PPL) nor recovered downstream capability above pretrained. Thus, we conclude that by neither route does the naive attacker reach the gain achieved by fine-tuning an unlocked model. Consistent with previous results in non-biological models, an informed attacker who implements the SVD-chain construction does recover capability on pathogenicity prediction, at the cost of increased computational requirements for the fine-tuning process. Availabilityhttps://github.com/Georgakopoulos-Soares-lab/glm-locking.

14
Reference-guided immune recovery matching prioritizes traditional Chinese medicine ingredients

Hu, C.; Xiao, B.; Chen, C. Y.-C.

2026-06-22 bioinformatics 10.64898/2026.06.16.732528 medRxiv
Top 1%
0.5%
Show abstract

Therapeutic prioritization from single-cell transcriptomes requires a target that is closer to treatment response than disease-signature reversal. In immune diseases, post-treatment recovery may follow patient- and cell-type-specific trajectories rather than a simple return along the pretreatment disease axis. We developed ImmuneNavi, a healthy-reference-anchored recovery-matching workflow for ranking traditional Chinese medicine ingredients from paired PBMC data. The workflow maps heterogeneous PBMC cohorts to a common healthy immune coordinate system, constructs patient-cell-type disease and recovery states, and processes ITCM treated-control profiles into a fixed ingredient perturbation bank. Patient and ingredient states are represented in matched gene, pathway and transcription-factor views, allowing the model to combine local transcriptional direction with more stable program-level features. A matcher trained on one paired treatment cohort preserved recovery-aligned ingredient rankings in independent PBMC cohorts without redefining the feature space, candidate set or preprocessing procedure. This provides a reusable transcriptomic pipeline for moving from paired immune-state measurements to prioritized natural-product candidates for experimental follow-up.

15
A high-level programming language for generative biology with Proto

Merchant, A.;Guo, D.;Viggiano, B.;Brennan-Almaraz, L.;Hurr, E.;Mai, T.;Yin, P.;King, S.;Ashley, E.;Hie, B.

2026-06-23 Synthetic Biology 10.64898/2026.06.22.733870 medRxiv
Top 1%
0.4%
Show abstract

Programmable composition of complex systems is a longstanding goal of biological research. Generative modeling has improved the reliability of computational design, but existing methods are highly specialized and are difficult to extend or compose. Here, we introduce Proto, a high-level programming language for generative biology. By composing a small set of abstract primitives into structured programs, Proto encodes generative design campaigns across diverse modalities and scales--spanning DNA, RNA, proteins, ligands, and their interactions. Proto readily incorporates predictive models into generative workflows, which we leveraged to design alternatively spliced introns with experimental validation in human cell lines. Proto is natively multi-objective, enabling the design of promoter-repressor pairs with leading experimental success rates for synthetic protein-DNA design. Alongside AI agents, Proto enables the specification of complex pathways and regulatory logic through natural language instructions. We openly release Proto, including software infrastructure and user interfaces, to enable widespread access to generative biological programming.

16
NLR from soybean Rsv1 locus confers broad-spectrum resistance to soybean mosaic virus G1-G7 strains by recognizing viral P3 protein

Zhao, H.; Gou, B.; Liao, J.; Zhao, Y.; Yang, T.; Huang, P.; Zhu, Y.; Tie, Y.; Wang, M.; Gao, L.; Li, K.; Zhi, H.; Cui, X.; Chen, X.; Xu, Y.; Duan, K.; Wang, Y.; Tao, X.

2026-07-09 plant biology 10.64898/2026.06.29.735421 medRxiv
Top 1%
0.4%
Show abstract

Nucleotide-binding leucine-rich repeat (NLR) immune receptor genes are of significant value in disease resistance breeding and the control of viral diseases. Soybean mosaic virus (SMV) poses a serious threat to soybean production and the Rsv1 locus in soybean cultivar Suweon 97 confers broad-spectrum resistance against SMV strains G1 to G7; however, this locus harbors no fewer than 18 NLR genes, and thus the broad-spectrum antiviral mechanisms underlying the Rsv1 locus remain poorly understood to date. Here, we established a rapid and highly efficient screening system for cloning NLR genes from soybean Rsv1 locus and identified a broad-spectrum antiviral NLR gene 13g184900 from this highly complicated locus. The NLR encoded by 13g184900 can recognize viral P3 protein from all SMV strains (G1-G7) and another potyvirus Bean common mosaic virus (BCMV). The coiled-coil (CC) domain of this NLR directly interacts with viral P3 protein. Additionally, we showed that this NLR originated from wild soybean accession in East China and has been introduced into several soybean cultivars during domestication. Collectively, we developed a high-throughput screening system for identifying NLR genes in soybean and our study provides new mechanistic perspective on how the Rsv1 locus mediates the broad-spectrum resistance to all SMV G1-G7 strains.

17
OCellus: A Language-Model Framework for Single-Cell, Spatial, and Perturbation Biology with Natural-Language Reasoning

Zhang, C.; Sun, J.; Xu, Z.; Liao, R.; Yin, A.; Gao, H.; Liu, E.; Bao, Y.; Zhao, L.; Wang, G.

2026-07-12 bioinformatics 10.64898/2026.07.08.737248 medRxiv
Top 1%
0.4%
Show abstract

Computational modeling of cellular behavior--the virtual cell--has emerged as a stated grand challenge at the intersection of artificial intelligence and biology, yet existing foundation models remain specialized: single-cell models process dissociated transcriptomes only, spatial models require dedicated spatial-aware architectures, and perturbation predictors depend on manually curated knowledge bases that cap generalization. Here we introduce OCellus, a single nine-billion-parameter language model (Qwen3.5-9B) fine-tuned on twenty-two biological tasks that simultaneously addresses all three limitations through three coordinated technical contributions on a shared backbone. First, EvenClock encodes two-dimensional spatial coordinates as eighteen clockface sectors of text, enabling spatial reasoning on a vanilla language model without architectural modification; on ten spatial transcriptomics tasks OCellus attains 77 percent spatial-neighborhood accuracy, 96 percent spatial-cellchat accuracy, and 0.70 proportion-cosine similarity on spatial deconvolution, all without any spatial-aware architectural components. Second, per-gene language-model embeddings replace the Gene Ontology annotations that GEARS depends on, achieving Pearson correlation 0.945 on the Replogle 2022 perturbation benchmark versus 0.84 for GEARS across 457 completely unseen knockout genes. Third, OCellus-Agent provides a Planner-Router-Verifier natural-language interface that achieves 75 percent pipeline accuracy on eighty multi-task queries. Removing language-model embeddings collapses perturbation Pearson to 0.06, confirming that learned functional representations--not graph topology--drive the gain. As a cell-type encoder, OCellus ranks first among fourteen foundation models in linear-probe accuracy at 95.1 percent across four benchmark datasets, and reaches 72.6 percent average across twenty-two evaluated biological tasks--a 57-percentage-point absolute gain over the strongest baseline configuration. As a language model, OCellus uniquely generates natural-language explanations of its predictions, a capability absent from all competing methods. Code, pre-trained model weights, the graph-neural-network module, and the agent system will be made available upon publication.

18
A geometric atlas of how ESM3 organizes modalities across depth

Steenwyk, J. L.

2026-07-12 bioinformatics 10.64898/2026.07.08.737319 medRxiv
Top 2%
0.4%
Show abstract

Protein language models learn general-purpose representations from large collections of protein sequences and structures, and have advanced the prediction of protein structure and function. ESM3 is a multimodal protein language model that ingests a protein through several channels at once, including amino-acid sequence, three-dimensional structure, secondary structure (SS8), solvent accessibility (SASA), and discrete functional annotations, summing their embeddings into a single residual stream. Little is known about whether these modalities occupy separate subspaces and the depth at which they fuse. The present analysis examines ESM3 (esm3-sm-open-v1; 1.4 billion parameters; 48 transformer layers) once per modality in isolation and applies representational-similarity analysis across all 48 layers. The four physical modalities (sequence, structure, SS8, SASA) begin in distinct subspaces, remain maximally separated through roughly the first half of layers, and then fuse into a shared low-dimensional subspace between layers 25 and 35. The fusion is ordered. The structure-derived modalities (structure, SS8, SASA) are mutually aligned from the input, whereas sequence joins last, after layer 28. The functional-annotation modality never fuses; instead, it remains representationally orthogonal to the physical modalities at every layer, and this orthogonality holds whether the annotation is supplied as whole-protein or per-residue, suggesting that it is content-driven rather than a tokenization arti-fact. The fusion is a learned property, absent in a randomly initialized model of the same architecture, holds at the residue level below the mean-pool, and reorganizes variance, converting between-condition variance into within-condition variance while the stream never approaches isotropy. Fusion depth is independent of protein length but is delayed by structural disorder. The phenomenon is universal across diverse organisms. Across 5,555 proteins from 12 organisms spanning eukaryota, bacteria, and archaea, every superkingdom (and every individual organism) reaches peak modality fusion at the same network depth (layer 35).

19
Genomic Surveillance of Respiratory Syncytial Virus among Patients with Acute Respiratory Infection through Hospital-Based Influenza Surveillance Platforms in Bangladesh, August 2024-December 2025

Alam, M. S.; Begum, M. N.; Rahman, M.; Chowdhury, F.; Jubair, M.; Karim, Y.; Shanto, M. R. R.; Howlader, R.; Rahman, T.; Talha, M.

2026-07-03 evolutionary biology 10.64898/2026.07.02.736091 medRxiv
Top 2%
0.4%
Show abstract

Background: RSV is a major cause of severe lung infections in young children, with over 95% of deaths occurring in poorer countries. Bangladesh has high rates of RSV illness in children but lacks genetic data from after the COVID-19 pandemic. New vaccines and antibody treatments are now available, making local genetic information essential. Objectives: We sequenced complete RSV genomes from Bangladeshi patients to study virus types, genetic changes, and protein mutations, and shared our data openly. Methods: From August 2024 to December 2025, we took 59 RSV-positive samples with high virus levels from hospital patients and sequenced their full genomes using Oxford Nanopore technology. Results: Among 11,874 patients, 1,390 (11.7%) had RSV, mostly RSV-A (94.6%). We obtained 49 good-quality full genomes from the 59 samples (83% success): 43 RSV-A (ON1 type, five sub-lineages) and 6 RSV-B (BA9 type). We found S276N in 35% of RSV-A and S389P in all RSV-B, but neither stops current antibody treatments. All RSV-A viruses gained a new sugar attachment site on their F protein, and most RSV-B viruses gained one too. We uploaded all 49 genomes to GISAID for public use. Conclusion: This work shows we can do full RSV genome sequencing in Bangladesh. The viruses here still match the targets of new vaccines and antibodies, which is reassuring. Our findings provide a foundation for planning RSV prevention in Bangladesh and South Asia.

20
CellPilot: an agentic framework that pilots small language models through autonomous single-cell annotation

Jiang, S.; Qi, C.; Chen, Y.; Song, X.; Wei, Z.

2026-07-10 bioinformatics 10.64898/2026.07.06.736807 medRxiv
Top 2%
0.4%
Show abstract

Large language models can annotate cell types from marker gene lists, but they typically operate after preprocessing and clustering are complete, treating annotation as a terminal labeling step rather than controlling the analytical decisions that produce the evidence for cell identity. We present CellPilot, an agentic framework that guides a locally deployable small language model through the full single-cell analysis workflow, from raw count matrices to cluster-level annotation. CellPilot combines standard single-cell analysis tools with structured workflow control and observation-guided reasoning, allowing the model to plan analyses, execute tools, inspect intermediate results and revise decisions within a traceable session. On GTEx, structured workflow orchestration raised the same 8B model from 0.39 in a prompt-only setting to 0.89, closing most of the gap to GPT-4o (0.92) within the same framework; the framework gain was substantially larger for the smaller backbone across datasets (+0.35 versus +0.19). Across GTEx, Tabula Sapiens, and Mouse Cell Atlas, CellPilot achieves cluster-level annotation accuracies of 0.891, 0.750, and 0.773, outperforming representative reference-based, marker-based, and LLM-based methods. CellPilot confidence scores were associated with annotation correctness and supported post hoc filtering, while complete execution traces were retained for each analysis. These results suggest that structured workflow orchestration can be a critical determinant of performance in multi-step single-cell analysis, enabling locally deployable small language models to approach larger proprietary models while preserving transparency and practical usability.